Back

DNA Research

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match DNA Research's content profile, based on 26 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Chromosome-scale genome assembly and annotation of the Vietnamese indica rice cultivar Khang Dan 18

Nguyen, T. Q.; Do, K. H. D.; Vu, T. M.; Hoang, N. V.

2026-08-21 plant biology 10.64898/2026.08.15.742683 medRxiv
Top 0.1%
6.7%
Show abstract

Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

2
Valeriana officinalis genome sequence reveals candidate genes for valerenic acid biosynthesis and flavonoid metabolism

de Oliveira, J. A. V. S.; Baez, M.; Pucker, B.

2026-08-21 genomics 10.64898/2026.08.14.744958 medRxiv
Top 0.1%
5.2%
Show abstract

Valeriana officinalis is the scientific name for valerian, a plant known for producing valerenic acid, a compound with anxiolytic properties. Anxiety disorders represent a significant global health crisis, impacting everyday lives. As the global demand for natural, non-synthetic anxiety treatments rises, V. officinalis has emerged as a promising, yet underutilized, medicinal resource. Understanding its genome is the first step toward unraveling the biosynthetic genes underlying valerenic acid production, facilitating further research into its production. Here, we report the first genome sequence of valerian, with an assembly size of 3.3 Gbp and an N50 of 110.8 Mbp, and its corresponding annotation with 96.6% completeness, providing a foundational resource for studying the genetic basis of specialized metabolism in valerian. The value of this genome sequence for discoveries in specialized metabolism is demonstrated by the identification of the flavonoid biosynthesis gene repertoire and the selection of strong candidate genes for valerenic acid biosynthesis. This genome sequence holds the potential to support future functional studies aimed at elucidating the regulation of medically relevant metabolite pathways in V. officinalis.

3
A subgenome-resolved and chromosome-scale reference genome assembly of allotetraploid wheat wild relative Aegilops peregrina

Singh, J.; Gudi, S.; Maughan, P. J.; Gill, U.; Gupta, R.

2026-08-30 genomics 10.64898/2026.08.28.747929 medRxiv
Top 0.1%
3.9%
Show abstract

Aegilops peregrina is a wild allotetraploid wheat wild relative and an important source of genetic diversity for stress tolerance and agronomic traits. Here, we report a subgenome-resolved, chromosome-scale reference genome assembly of a drought tolerant and stem rust resistant Ae. peregrina accession PI 604178 generated using PacBio HiFi and Hi-C sequencing. The 10.13 Gb assembly contains 98.81% of sequence anchored to 14 pseudomolecules representing the seven S and seven U chromosomes, with contig and scaffold N50 values of 25.84 and 746.48 Mb, respectively. The assembly achieved a consensus quality value of 74.61, 97.83% k-mers completeness, and 99.9% BUSCO completeness. LTR Assembly Index values of 20.43 and 18.79 for the S and U subgenomes, respectively, further supported high continuity across repeat-rich regions. Repetitive elements comprise 85.93% of chromosome-anchored assembly. We annotated 59,910 high-confidence protein-coding genes, with comparable gene representation across the two subgenomes. This reference genome provides a high-quality genomic framework for comparative analyses, characterization of important loci regulating agronomic and resilience related traits, and sequence-guided exploitation of Ae. peregrina allelic diversity for wheat improvement.

4
Whole genome sequences and annotations of Japanese and French strains of Heterosigma akashiwo

Kondo, T.; Sakamoto, M.; Tokumaru, M.; Tanizawa, Y.; Nakamura, Y.; Toyoda, A.; Ueki, S.

2026-08-23 genomics 10.64898/2026.08.19.745619 medRxiv
Top 0.1%
2.7%
Show abstract

High-quality reference genomes provide an essential foundation for elucidating the molecular basis of organismal ecophysiology. Here, we sequenced and assembled chromosome-scale genomes of two Heterosigma akashiwo strains isolated from coastal waters of Japan and France. The assembly sizes were 1.18 Gb and 1.43 Gb for the Japanese and French strains, respectively. The scaffold N50 of the Japanese strain assembly was 66 Mb, whereas the one of the unscaffolded French strain assembly was 33 Mb. To our knowledge, these assemblies represent among the largest and most contiguous genome resources currently available for members of the Stramenopiles (Ochrophyta). Evidence-based gene prediction in the Japanese strain recovered approximately 90% of conserved stramenopile core genes, indicating a highly complete gene repertoire, and was complemented by extensive functional annotation. In the French strain, homology-based gene prediction recovered approximately 80% of conserved core genes. Comparative genome analysis revealed extensive synteny conservation between the two strains, although several putative duplication and translocation events were detected. These genomic resources provide a robust framework for investigating the molecular, cellular, and ecological mechanisms underlying the physiology, adaptation, and bloom-forming capacity of H. akashiwo.

5
Chromosome-level genome assemblies and annotations of Amaranthus spinosus, Amaranthus acanthochiton, Amaranthus arenicola, and Amaranthus floridanus

Raiyemo, D. A.; Werle Noe, I.; Kaur, R.; Whitt, L.; Carey, S. B.; Hale, H.; Lewis, K. J.; Womack, L.; Harkess, A.; Llaca, V.; Fengler, K.; Patterson, E. L.; Gaines, T. A.; Tranel, P. J.

2026-08-25 genomics 10.64898/2026.08.21.746229 medRxiv
Top 0.1%
2.4%
Show abstract

Amaranthus L. spans aggressive agricultural weeds, ornamentals, and ancient pseudocereals. Species within the genus vary in morphology, environmental tolerance, and sexual systems, making them well-suited for studying reproductive evolution and plant adaptation. To investigate sex chromosome architecture within the genus, we generated chromosome-level assemblies of a monoecious amaranth (Amaranthus spinosus) and three dioecious species (A. acanthochiton, A. arenicola, and A. floridanus) using PacBio high-fidelity (HiFi) long reads. We paired these data with Dovetail Genomics Omni-C sequencing to achieve haplotype resolution for A. spinosus and A. acanthochiton, and we used reference-guided scaffolding for the remaining two species. The assemblies are highly contiguous, with sizes ranging from 394.24 to 607.10 Mbp, contig N50 from 0.63 to 8.76 Mbp, and scaffold N50 from 22.44 to 37.97 Mbp. Evaluation of the assemblies and annotations revealed 96.3 to 97.6%, and 97.6 to 98.3% BUSCO completeness, respectively. Comparative genomic analysis revealed that the Chromosome 1 inversions and Robertsonian fusion previously reported in A. tuberculatus are conserved in A. acanthochiton and consistent with the architecture of A. arenicola and A. floridanus, suggesting that the evolution of dioecy in this clade predates subsequent speciation. In parallel, multiple homologs of Rf1 on Chromosome 3 of A. spinosus, a monoecious species that exhibits spatial separation of male and female flowers and is closely related to the dioecious A. palmeri, were identified. Together, this study provides foundational resources for advancing evolutionary, ecological, and agronomic research across the genus, including herbicide resistance evolution and weediness traits.

6
Chromosome-level genome assembly of the European leaf-toed gecko, Euleptes europaea

Paris, J. R.; Abueg, L.; Pelan, S.; Sims, Y.; Tilley, T.; Mountcastle, J.; Balacco, J.; OToole, B.; Fedrigo, O.; Formenti, G.; Jarvis, E. D.; Canestrelli, D.; Salvi, D.

2026-08-18 genomics 10.64898/2026.08.10.744031 medRxiv
Top 0.1%
2.4%
Show abstract

The European leaf-toed gecko (Euleptes europaea) is a small, nocturnal gecko endemic to the western Mediterranean. As a phylogenetically distinctive member of the Gondwanan family Sphaerodactylidae, it represents an important species for studying Mediterranean island biogeography, adaptation, and reptile genome evolution. The species also occupies a key position for investigating the evolution of sex chromosomes, as geckos exhibit remarkable diversity and frequent transitions in sex-determination systems. We present a chromosome-level genome assembly of Euleptes europaea generated as part of the Vertebrate Genomes Project. The 1.8 Gb assembly has a scaffold N50 of 102.3 Mb (contig N50 27 Mb), with 21 chromosome-scale scaffolds corresponding to the known karyotype (2n = 42). The primary assembly has a BUSCO completeness of 97.80% (95.60% as single-copy), a k-mer completeness of 96.00%, and a k-mer quality value (QV) of 61.20. Repetitive elements account for 53.20% of the genome and genome annotation identified 18,633 protein-coding genes. This high-quality reference genome will facilitate studies of genome evolution, island adaptation, and sex chromosome evolution across geckos and other reptiles.

7
Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.

2026-08-29 genomics 10.64898/2026.08.26.747264 medRxiv
Top 0.2%
2.0%
Show abstract

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

8
Gene model for the ortholog of DENR in Drosophila pseudoobscura

Lawson, M. E.; Sanow, K.; Fratian, M.; Matura, M.; Scanlon, R.; Richard, M.; Nakhla, M.; Rele, C. P.; Thompson, J. S.; Findlay, G. D.; O'Rourke, K. S.

2026-08-11 genomics 10.64898/2026.08.11.744233 medRxiv
Top 0.2%
1.3%
Show abstract

Gene model for the ortholog of Density regulated protein (DENR) in the Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

9
Haplotype-resolved chromosome-level genome assembly of four European white oak species

Magris, G.; Avanzi, C.; Bagnoli, F.; Duvaux, L.; Belmonte, E.; Vendramin, G. G.; Piotti, A.; Pinosio, S.

2026-08-24 genomics 10.64898/2026.08.20.745905 medRxiv
Top 0.2%
1.1%
Show abstract

European white oaks (Quercus section Quercus) are ecologically and economically important forest trees characterized by extensive shared genetic variation and a history of interspecific gene flow. Genomic resources remain uneven across species, limiting comparative analyses and pangenome development. Here, we present haplotype-resolved chromosome-scale genome assemblies and genome annotations for four European white oak species: Quercus robur, Q. petraea, Q. pubescens, and Q. frainetto. The assemblies were generated from PacBio HiFi sequencing data and include both phased haplotypes for each species. Genome sizes range from 779 to 817 Mb and all assemblies are organized into 12 chromosome-scale pseudomolecules with high completeness and contiguity. We additionally provide species-specific repeat annotations, structurally and functionally annotated protein-coding gene sets, and complete organellar genomes. The dataset includes the first reference genomes for Q. pubescens and Q. frainetto, together with newly generated assemblies for Q. robur and Q. petraea produced using a consistent sequencing and analysis workflow. These resources provide a standardized framework for comparative genomics, pangenome construction, genome evolution studies, and investigations of adaptation and introgression across European white oaks.

10
Gene model for the ortholog of Ilp3 in Drosophila pseudoobscura

Lieser, B. C.; Laskowski, L. F.; Huber, R.; Kolker, K. O.; Arsham, A. M.; Rele, C. P.; Toering Peters, S.

2026-08-23 genomics 10.64898/2026.08.19.745830 medRxiv
Top 0.3%
1.1%
Show abstract

Gene model for the ortholog of Insulin-like peptide 3 (Ilp3) in the D. pseudoobscura Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

11
Chromosome assembly for the Black bean aphid Aphis fabae

Whitehead, M. A.; Claudia Wierzbicki, C.; Hughes, M.; Darby, A. C.

2026-08-11 genomics 10.64898/2026.08.05.743085 medRxiv
Top 0.3%
1.0%
Show abstract

The black bean aphid, Aphis fabae is a crop pest and vector of insect-transmitted pathogens, comprising closely related sub-species with overlapping host ranges. In other Aphis species, over-expression of specific detoxification genes has been linked to insecticide tolerance. We present two chromosome-scale assemblies for a clonal A. fabae line, representing two phased haplotypes, generated using HiFi and Hi-C sequencing technologies. A comprehensive genome annotation, built with PacBio Iso-Seq data, was used to investigate genes underlying insecticide tolerance. Both genomes are comprised of four chromosomal blocks (haplotype 1: 427 Mb; haplotype 2: 396 Mb) with high BUSCO completeness (98.7%). Comparative genomics revealed an expansion of UDP-glycosyltransferases, whose expression is linked to insecticide detoxification in other Aphis species. These high-quality references provide a foundation for studying A. fabae sub-species and a genomic resource for investigating insecticide tolerance across the Aphis genus. Author summaryHere we have provided a comprehensive assembly and annotation for further study into the Black bean aphid, Aphis fabae, using up to date long-range sequencing technologies. The final assemblies for both haplotypes are chromosome length and consist of 4 main chromosome blocks, consistent with the literature. The A. fabae genome was found to contain an increase in copy number of UDP-glycosyltransferases, which have previously been linked to insecticide resistance. The work here will be a resource to those studying insecticide tolerance in crop pests, as well as the differences between A. fabae sub-species.

12
UPLC-ESI-MS based lipidomics revealed novel biomarkers in insulin receptor knockdown induced type 2 diabetes model of Drosophila

Kumar, P.; Fatima, Z.; Kumar, P.; Kumar, R.; Chauhan, B. S.; SRIKRISHNA, S.

2026-08-20 biochemistry 10.64898/2026.08.20.745875 medRxiv
Top 0.3%
0.9%
Show abstract

Type 2 diabetes (T2D) is a prevalent metabolic disorder affecting millions worldwide, characterized by insulin resistance and impaired glucose homeostasis. While mammalian models are widely used, Drosophila melanogaster provides a powerful alternative due to its conserved insulin signaling pathways, genetic tractability, and suitability for high throughput studies. In addition to glucose dysregulation, lipid metabolism plays a crucial role in T2D pathophysiology, as alterations in lipid composition contribute to insulin resistance and metabolic dysfunction. Lipidomic studies have emerged as an essential approach to identify metabolic signatures and potential biomarkers for disease progression and therapeutic targeting. In this study, T2D like model was established by inducing insulin resistance through knockdown of the insulin receptor in brain insulin-producing cells using the dilp2-Gal4>UAS-InRRNAi system. This genetic manipulation resulted in significant metabolic dysregulation, including elevated glucose, trehalose, and triacylglyceride levels, along with increased oxidative stress indicators. Additionally, mRNA expression analysis of key insulin signaling components, including insulin receptor substrate 1, dilp2, dilp3, dilp5, and phosphorylated Akt, further validated the model. To further investigate metabolic alterations, Lipid profiling was performed using ultra-performance liquid chromatography coupled with quadrupole time-of-flight mass spectrometry (UPLC-QTOF-MS) in non targeted LC-MS-based metabolomics approach to identify lipid biomarkers associated with T2D. Multivariate statistical analyses, including PCA and PLS-DA, revealed distinct lipid signatures between wild-type and T2D flies. Notably, specific phosphatidylglycerol species PG 34:0, PG 34:4, PA 38:3, PIP 38:1, PIP2 38:6, and LPS 24:0 demonstrated an area under the curve (AUC) of 1, indicating their strong reliability as lipid biomarkers for T2D diagnosis.

13
Comprehensive analysis of mulberry genetic diversity based on 1-DNJ content and SNP markers

Shen, Z.; Li, J.; Shi, J.; Li, Z.; Wang, F.; Geng, J.; Hu, K.

2026-08-19 genetics 10.64898/2026.08.11.744330 medRxiv
Top 0.3%
0.9%
Show abstract

Mulberry trees have high economic and ecological value, and a robust molecular marker system plus germplasm genetic diversity analysis is critical for innovative utilization of high-quality medicinal and economic mulberry germplasm. Here, 51 mulberry samples were used to develop SNP primers via genome resequencing, with the SNP-PCR system optimized by single-factor and orthogonal assays. The phenotypic diversity and SNP molecular marker genetic diversity of 1-deoxynojirimycin (1-DNJ) in mulberry leaves were analyzed respectively, and the genetic correlation between molecular markers and phenotypic traits was evaluated by Mantel test. Tested germplasm showed marked 1-DNJ variation (0.4805-2.5300 mg/g, CV=0.4241), reflecting rich genetic diversity. The optimal SNP-PCR system included Buffer (containing Mg{superscript 2}+) 2.2 L, 2.5 mM dNTP 0.4 L, forward and reverse primers (10 mol{middle dot}L-1) totaling 2.75 L, Taq DNA polymerase (5 U{middle dot}L-1) 0.3 L, DNA (50 ng{middle dot}L-1) 1.1 L, and ddH2O 13.65 L. 23 highly polymorphic ones amplified 91 loci (81 polymorphic, 89.10% polymorphism rate). Genetic diversity analysis showed that the average genetic distance was 0.3010, and the average expected heterozygosity (H) and Shannon information index (I) reached 0.4667 and 0.3104 respectively, indicating that the genetic differentiation among the tested mulberry germplasms was significant and the population had a moderate to upper level of genetic diversity. UPGMA clustering divided 51 germplasms into 6 major groups at a genetic similarity coefficient of about 0.7, while phenotypic clustering based on 1-DNJ content divided them into 2 major categories and 4 subcategories, with high 1-DNJ germplasm clustered independently. Mantel correlation analysis showed that 6 SNP sites were significantly weakly correlated with 1-DNJ content (r < 0.3, p < 0.05), and can be used as candidate molecular markers for subsequent genetic analysis of 1-DNJ content.This study established a stable mulberry SNP-PCR system, Analyze the molecular genetic characteristics of mulberry germplasm and DNJ phenotypic variation rules respectively, and provide basic data for cluster comparison. and provided a scientific basis for marker database improvement, germplasm identification and molecular-assisted breeding.

14
Rapid evolution and functional divergence of the monkeyflower Mimulus lewisii telomerase

Samo, N.; Nguyen, L.; Kumawat, S.; Choi, J. Y.

2026-08-09 evolutionary biology 10.64898/2026.08.05.739867 medRxiv
Top 0.3%
0.8%
Show abstract

Telomeres are nucleoprotein structures that protect chromosome ends and are maintained by the Telomerase Reverse Transcriptase (TERT) protein that uses a noncoding Telomerase RNA (TR) as a template. In monkeyflowers, Mimulus lewisii had an ancient TR gene duplication, synthesizing an evolutionarily atypical sequence heterogeneous telomere. How TERT interacts with both TR paralogs during telomere maintenance is unknown and answers can shed novel insights underlying telomere function. Using new genome assemblies we discovered TERT is rapidly evolving in lineages sharing the TR duplication. We investigated the functional consequences arising from the rapid evolution, first by using yeast three-hybrid and testing the physical binding between conspecific and heterospecific TERT-TR combinations. Results showed TERT binds both ancestral (TR1) and derived (TR2) TR paralogs in M. lewisii, but not in species without a functioning TR2. We located the region of TR binding to amino acids near the KRxR motif. We then combined next-generation sequencing with Telomeric Repeat Amplification Protocol and discovered M. lewisii had high telomerase activity. Comparative transcriptomics indicated no strong evidence of expression divergence in telomere maintenance genes for M. lewisii, suggesting rapid evolution shaped TERT protein sequence. In vivo activity of M. lewisii telomerase was investigated by analyzing F1 telomeres generated by crossing M. lewisii and M. verbenaceus, which doesnt have a functioning TR2. Results showed M. verbenaceus chromosome ends in the F1 had converted into M. lewisii telomeres, suggesting dominance of the M. lewisii telomerase. We demonstrate TERT-TR coevolution can have significant consequences on the evolution of plant telomeres. Significance statementTelomeres protect chromosome ends and are maintained by the telomerase complex. We discovered the catalytic component of the telomerase (TERT) was rapidly evolving in monkeyflowers (Mimulus) and studied the molecular consequences. In M. lewisii, TERT evolved lineage-specific amino acids to bind two sequence divergent telomerase RNA paralogs. Telomerase activity assay showed M. lewisii synthesized more telomere repeats compared to its sister species without the TR duplication, and transcriptomics indicated this was not due to a change in telomere maintenance gene expression. Genetic experiments in interspecies hybrids showed M. lewisii telomerase could convert chromosome ends in sister species into M. lewisii-like telomeres suggesting functional dominance. We show rapid evolution of the telomerase can have significant effects on telomere evolution.

15
Chloroplast Genome Evolution, Heteroplasmy, and Inverted Repeat Dynamics in the Elymus Complex (Triticeae, Poaceae): Insights from Single-Molecule Sequencing of Elymus ciliaris and Comparative Analysis of St-Genome Lineages

Karimi, N.; Zhang, Y.; Saeidi, H.; Schwarzacher, T.; Liu, Q.; Heslop-Harrison, J. S.

2026-08-07 genomics 10.64898/2026.08.03.742459 medRxiv
Top 0.4%
0.8%
Show abstract

Background/ObjectivesElymus sensu lato (Poaceae) is arguably the largest and most complex genus in the tribe Triticeae. It includes hybrids and polyploids based on x=7 chromosomes, all including the St genome, forming a valuable genepool for forage grass and cereal breeding. Analysis of chloroplast genome diversity and structural dynamics is critical for resolving maternal lineages, reticulate evolution and biodiversity across this agronomically important complex, refining their taxonomy, conservation and exploitation. MethodsWe sequenced the complete chloroplast genome (plastome) of Elymus ciliaris (4x=2n=28; StStYY genome composition) using ultra-long Oxford Nanopore single-molecule reads and compared it to 76 additional chloroplast genomes representing major St-genome lineages in Elymus s.l. (Pseudoroegneria St; Elymus s.s. StH, StY; Thinopyrum StJ/E; Campeiostachys StYH; Kengyilia StYP). We analyzed structure, nucleotide diversity, inverted repeat (IR) dynamics, and phylogenetic signal. ResultsThe E. ciliaris chloroplast genome was 135,004 bp long (38.3% GC) with a canonical quadripartite structure. Single-molecule reads (n=74) revealed heteroplasmy: two Small-Single-Copy (SSC) orientations at 30%:70% frequency, indicating an inversion polymorphism. Across the Elymus group, comparative analysis of chloroplast assemblies showed high structural conservation but lineage-specific IR-boundary shifts. Kengyilia exhibits exceptional IR expansion. Nucleotide diversity hotspots localize to the large single-copy region, especially in StY lineages. Phylogenies recover a monophyletic St-containing clade but do not delineate genera, reflecting reticulate evolution, with North American/Southeast Asian and Eurasian geographic sub-clades. ConclusionsSingle-molecule sequencing uncovered heteroplasmy with an inversion polymorphism in a single plant of Elymus ciliaris, hidden in short read assemblies. There were no other polymorphisms, as expected for chloroplast sequences (except for technical homopolymer variation). Our analyses showed that a Pseudoroegneria-like St chloroplast genome predominates as the maternal donor across Elymus polyploids. Variable regions and IR dynamics offer strong models for chloroplast genome evolution in reticulate lineages and suggest exploiting plastome variation to complement nuclear biodiversity studies.

16
A chromosome-scale genome assembly of the Swiss Lolium multiflorum ecotype Tremona reveals a scalable method to purge spurious duplications

Piat, L.; Herren, G.; Grieder, C.; Roulin, A. C.

2026-08-20 genomics 10.64898/2026.08.18.745395 medRxiv
Top 0.4%
0.8%
Show abstract

Italian ryegrass (Lolium multiflorum) is a key temperate forage species underpinning livestock production in Europe. Genomic resources remain limited by its large (2.2 Gb), repetitive, and highly heterozygous genome. Here, we present a high-quality chromosome-scale genome assembly of the Swiss L. multiflorum ecotype Tremona, collected in 2008 in Ticino, Switzerland, and subsequently incorporated into recurrent breeding cycles in the Swiss breeding program. To address systematic assembly artefacts caused by unresolved haplotypes in our initial PacBio HiFi assembly, we developed ParaLies, a post-assembly tool that identifies and removes artefactual duplications based on sequence divergence while preserving true paralogous gene copies. ParaLies reduced the duplicated BUSCO rate from 16.91% to 6.72% without loss of bona fide genomic content. The resulting assembly has a contig N50 of 15.69 Mb and captures 94% of the expected 2.2-Gb genome size. We further analyzed whole-genome resequencing data from Tremona, additional Swiss ecotypes, and publicly available North American germplasm. Tremona was genetically homogeneous, with no evidence of pronounced recent bottlenecks or substantial within-population structure, and was genetically distinct from the other Swiss ecotypes analyzed. Together, the Tremona genome and ParaLies provide valuable resources for L. multiflorum genomics and breeding and demonstrate a scalable approach for reducing haplotype-induced redundancy in highly heterozygous genomes.

17
Prime-Editing in Marchantia paleacea: Expanding the Genome-Editing Toolbox in Bryophytes

Danilo, B.; Quillien, A.; Rojas-Latorre, C.; Nibani, Z.; Mestre, C.; Delaux, P.-M.; Lauressergues, D.; Neveu, J.

2026-08-11 plant biology 10.64898/2026.08.07.743462 medRxiv
Top 0.4%
0.6%
Show abstract

Since the development of CRISPR-based genome editing tools, a number of novel technologies have emerged. This includes Prime-Editing that acts as a search and replace genome editing tool. Prime-Editing has been deployed across multiple clades, including in a few flowering plants. Here, we report on the development of an efficient Prime Editor (PE) for the model bryophyte Marchantia. Initial tests were conducted on Acetolactate Synthase as a target and revealed an average efficiency above 40%. The system has been developed in the GoldenGate cloning system, facilitating construct design. The development of PE in Marchantia expands the Genome-Editing tools available for this emerging model in plant biology.

18
A naturally occurring frameshift mutation in the UNUSUAL FLORAL ORGANS gene associated with the marimo floral phenotype in gerbera

Hattori, T.; Shimada, R.; Nagakura, M.; Ando, R.; Isobe, S.; Tajima, N.; Hirakawa, H.; Shirasawa, K.; Tominaga, A.

2026-08-14 genetics 10.64898/2026.08.09.743735 medRxiv
Top 0.5%
0.5%
Show abstract

BackgroundThe capitulum of Asteraceae is a highly specialized inflorescence whose formation requires the coordinated regulation of multiple developmental processes, including floral organ identity and floral meristem determinacy. The LEAFY (LFY)-UNUSUAL FLORAL ORGANS (UFO) regulatory module is known to play an important role in flower development; however, naturally occurring mutations affecting this pathway have not been genetically characterized in gerbera (Gerbera hybrida). ResultsIn this study, we characterized a novel gerbera mutant identified during a commercial crossing program and named it marimo based on its green, spherical capitulum. Morphological observations revealed the repeated formation of secondary and tertiary floret-like organs within primary floret-like organs. Scanning electron microscopy showed that the epidermal structure of the green organs in marimo was similar to that of wild-type involucral bracts. RNA sequencing identified numerous differentially expressed genes between marimo and the wild type, and network and Gene Ontology analyses highlighted gene groups associated with flower development, reproductive organ differentiation, and tissue structure formation. RNA-seq analysis showed increased expression of LFY and reduced expression of GGLO1, a PISTILLATA/GLOBOSA-like B-class MADS-box gene, in the marimo mutant. RT-qPCR analysis of a segregating population further confirmed reduced GGLO1 expression in marimo-type individuals. In addition, a single-nucleotide deletion was identified in the coding region of UFO. This deletion was predicted to cause a frameshift and a premature stop codon. In selfed progeny of No. 251, the UFO genotype was fully associated with capitulum phenotype, and only individuals homozygous for the mutant allele exhibited the marimo phenotype. ConclusionsThese results indicate that the naturally occurring frameshift mutation in UFO is the strongest candidate variant underlying the marimo phenotype. RNA-seq analysis showed increased LFY expression and markedly reduced GGLO1 expression in the marimo mutant. Reduced activity of the LFY-UFO regulatory module may therefore have altered the expression of GGLO1 and other floral organ development-related genes despite the continued expression of LFY. These changes may have affected both floral organ identity and floral meristem determinacy, resulting in the formation of green involucral bract-like organs and the repeated production of floret-like organs. The marimo mutant provides a useful genetic resource for investigating capitulum development in Asteraceae and may also serve as breeding material for introducing novel ornamental traits into gerbera.

19
Oncogenes have the most distinct codon biases in the genome and codon signatures that oppose tumor suppressor genes

Mathur, C.; Davis, E. T.; Ehrbar, D.; Omeoga, H. C.; Endres, L.; Byrne, S. R.; Begley, U.; Dedon, P. C.; Begley, T. J.

2026-08-24 cancer biology 10.64898/2026.08.21.746283 medRxiv
Top 0.5%
0.5%
Show abstract

Oncogenes and tumor-suppressor genes play opposing roles in cancer biology to promote and restrict growth, respectively. Codon usage patterns interface with tRNA modifications to control translation, leading to gene-specific codon signatures with regulatory potential. As such, codon-biased translational regulation has been identified as a driver of proliferation and drug resistance in multiple cancers. We used advanced codon analytics methods to characterize and compare codon usage bias in oncogenes and tumor suppressor genes (TSGs) from humans and mice at group and gene-specific levels. We demonstrate that human oncogenes exhibit a distinct and opposing codon usage pattern to TSGs. This phenomenon is also present in mice but with less distinct oncogene bias relative to humans. Further comparison to 447 gene ontology groups demonstrated that human oncogenes have the most distinct codon usage patterns in the genome, while also highlighting that codon bias can separate functionally related genes and pathways from other biological processes. Using gene-specific codon analytics, we determined that human oncogenes have two types of extreme codon bias: a large group (N = 43) over-using G/C ending (GC3) codons and a smaller group (N = 12) over-using A/U (AU3) ending codons. While GC3 bias has been linked to increased translation in general, the AU3 finding suggests that genetic, environmental, or stress-related signals could drive the translation of this small group of oncogenes. The less extreme bias observed in mouse oncogenes and tumor suppressors likely underscores species-specific differences in oncogenic translation programs. Together, our findings highlight codon usage bias as a potential determinant of oncogene expression, provide a framework for ontology-based codon analysis, and uncover on species-specific differences in oncogene translation and codon usage biases.

20
CannSelect: A High-Quality Genotyping Platform for Cannabis sativa

Wilkerson, D. G.; Stack, G. M.; Carlson, C. H.; Quade, M. A.; Dowling, C. A.; Toth, J. A.; Murdock, M. J.; Jasinski, J.; Stansell, Z. J.; McKay, J. K.; Smart, L. B.

2026-08-21 genomics 10.64898/2026.08.18.745408 medRxiv
Top 0.6%
0.4%
Show abstract

The field of genomics has enabled extraordinary progress in horticultural crop research. However, there is still a need for cost-effective, high-resolution technologies flexible to the diversity found in emerging crops. To this end, we introduce CannSelect, a high-quality genotyping platform for Cannabis sativa. Designed for use in diversity analyses and trait mapping, probe targets were selected from four genotyped diversity panels and a curated gene list. This platform has been used to effectively map day-neutrality in a segregating population to the Autoflower1 locus with average capture efficiencies of 88.5%. With broad genome coverage, demonstrated target specificity, and reproducibility, CannSelect is expected to perform well across the diversity of C. sativa. We describe the methodology used to design CannSelect v1.0 and performance metrics for testing capture efficiency and target alignment in diverse genome assemblies. The CannSelect platform represents a robust and scalable, genome-wide genotyping tool for C. sativa researchers and breeders.